Research & policy
Academic papers, official research, regulatory material, patents, and standards are grouped together with their evidence labels intact.
arXiv
Aug 04, 2026
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
Weekly significanceFor AI-product teams and investors, increasing inference spend is not a comparable performance lever unless the search procedure, verifier, compute accounting, and uncertainty reporting are specified. The paper offers a framework for separating real system improvements from reporting artifacts; its conclusions remain preprint evidence.
arXiv
Aug 03, 2026
In-Network Market Prediction Using Machine Learning and Limit Order Books
Weekly significanceThis is an early systems-research result, not evidence of trading profitability. It is relevant because it frames ultra-low-latency market prediction as a joint model-and-network-design problem and supplies a concrete offload trade-off for electronic-market infrastructure.
ECB research
Aug 06, 2026
Financial frictions across the production network and the transmission of monetary policy
Weekly significanceFor macro and credit-risk analysis, the result argues for tracking the location of leverage and financing stress across supply chains rather than relying only on aggregate leverage. The finding is a working-paper result and should be tested against other geographies and specifications before operational use.
arXiv cs.AI
Aug 07, 2026
FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
Weekly significanceFinancial AI systems can produce a plausible answer while citing the wrong filing or reporting period. FinRank makes provenance-sensitive retrieval measurable; its reported baselines show that this remains difficult. The results are author-reported preprint evidence, not a production-system validation.
arXiv
Aug 04, 2026
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
Weekly significanceModel routers and cost-quality cascades often lose latency savings when a handoff forces a new prefill. This is a promising systems result for serving economics, but the results are model-pair-specific preprint evidence—not a validated production saving across vendors or workloads.
arXiv cs.CL
Aug 06, 2026
The Bitter Lesson of Tool Calling
Weekly significanceTool-interface design can materially affect agent reliability and parallel-task performance without changing model weights. Teams should treat the result as a prompt to test typed, code-mediated tools in their own environment rather than as a general performance guarantee.
arXiv
Aug 05, 2026
Local Violation Certification for Linear Predict-Then-Optimize Pipelines
Weekly significanceAutomated pricing, allocation, and risk workflows increasingly couple machine learning with optimization. This is an early research result, but its framing of auditable local failure risk is relevant to financial and regulated decision systems where rare violations are costly.
arXiv
Aug 03, 2026
Why Large Language Models Fail at Tabular Prediction
Weekly significanceFor financial and enterprise analytics teams, the paper is a useful caution against treating a general LLM as a drop-in tabular model. The conclusion is a research finding from the authors' experimental setup, not a universal performance verdict for tool-augmented, fine-tuned, or purpose-built tabular systems.
arXiv cs.AI
Aug 06, 2026
Learning When to Trust via Selective Context Preference Optimization
Weekly significanceAgent and retrieval systems often act on external signals that may be stale, adversarial, or simply wrong. A selective-trust evaluation could be more decision-relevant than aggregate answer accuracy, but the result needs independent replication and production testing.
arXiv cs.AI
Aug 06, 2026
AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games
Weekly significanceThe cost of rigorous agent evaluation is becoming a bottleneck for model deployment and investment diligence. The method is a promising route to auditable early stopping, but its reported efficiency comes from a specific imperfect-information-game setting and is not yet a general LLM-evaluation result.